[ExecuTorch][WebGPU] Add shape/copy ops (dim_order, expand_copy, fill) to the WebGPU backend#20929
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/20929
Note: Links to docs will display an error until the docs builds have been completed. ❌ 4 New Failures, 1 Unrelated Failure, 6 Unclassified FailuresAs of commit a206e8b with merge base 21554e5 ( NEW FAILURES - The following jobs have failed:
UNCLASSIFIED FAILURES - DrCI could not classify the following jobs because the workflow did not run on the merge base. The failures may be pre-existing on trunk or introduced by this PR:
FLAKY - The following job failed but was likely due to flakiness present on trunk:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This PR needs a
|
SS-JIA
left a comment
There was a problem hiding this comment.
Review automatically exported from Phabricator review in Meta.
…) to the WebGPU backend Pull Request resolved: #20929 Add `_clone_dim_order` (DMA copy), `aten.expand_copy` (broadcast), and `fill` (`aten.full`/`full_like`/`scalar_tensor`) handlers for the training tail. Key changes: - `runtime/ops/{dim_order,expand_copy,fill}/` — WGSL kernels + handlers (dim_order is a native buffer DMA copy) - `CMakeLists.txt` `WEBGPU_SRCS` — wire the sources Reuses the shared Vulkan partitioner (`_clone_dim_order`/`expand_copy`/`full`* already registered); WebGPU kernels only. Co-authored-with: Claude Code. ghstack-source-id: 405026073 @exported-using-ghexport Differential Revision: [D111755131](https://our.internmc.facebook.com/intern/diff/D111755131/)
…) to the WebGPU backend Pull Request resolved: #20929 Add `_clone_dim_order` (DMA copy), `aten.expand_copy` (broadcast), and `fill` (`aten.full`/`full_like`/`scalar_tensor`) handlers for the training tail. Key changes: - `runtime/ops/{dim_order,expand_copy,fill}/` — WGSL kernels + handlers (dim_order is a native buffer DMA copy) - `CMakeLists.txt` `WEBGPU_SRCS` — wire the sources Reuses the shared Vulkan partitioner (`_clone_dim_order`/`expand_copy`/`full`* already registered); WebGPU kernels only. Co-authored-with: Claude Code. ghstack-source-id: 405026073 @exported-using-ghexport Differential Revision: [D111755131](https://our.internmc.facebook.com/intern/diff/D111755131/)
…) to the WebGPU backend Pull Request resolved: #20929 Add `_clone_dim_order` (DMA copy), `aten.expand_copy` (broadcast), and `fill` (`aten.full`/`full_like`/`scalar_tensor`) handlers for the training tail. Key changes: - `runtime/ops/{dim_order,expand_copy,fill}/` — WGSL kernels + handlers (dim_order is a native buffer DMA copy) - `CMakeLists.txt` `WEBGPU_SRCS` — wire the sources Reuses the shared Vulkan partitioner (`_clone_dim_order`/`expand_copy`/`full`* already registered); WebGPU kernels only. Co-authored-with: Claude Code. ghstack-source-id: 405026073 @exported-using-ghexport Differential Revision: [D111755131](https://our.internmc.facebook.com/intern/diff/D111755131/)
…) to the WebGPU backend Pull Request resolved: #20929 Add `_clone_dim_order` (DMA copy), `aten.expand_copy` (broadcast), and `fill` (`aten.full`/`full_like`/`scalar_tensor`) handlers for the training tail. Key changes: - `runtime/ops/{dim_order,expand_copy,fill}/` — WGSL kernels + handlers (dim_order is a native buffer DMA copy) - `CMakeLists.txt` `WEBGPU_SRCS` — wire the sources Reuses the shared Vulkan partitioner (`_clone_dim_order`/`expand_copy`/`full`* already registered); WebGPU kernels only. Co-authored-with: Claude Code. ghstack-source-id: 405026073 @exported-using-ghexport Differential Revision: [D111755131](https://our.internmc.facebook.com/intern/diff/D111755131/)
…) to the WebGPU backend Pull Request resolved: #20929 Add `_clone_dim_order` (DMA copy), `aten.expand_copy` (broadcast), and `fill` (`aten.full`/`full_like`/`scalar_tensor`) handlers for the training tail. Key changes: - `runtime/ops/{dim_order,expand_copy,fill}/` — WGSL kernels + handlers (dim_order is a native buffer DMA copy) - `CMakeLists.txt` `WEBGPU_SRCS` — wire the sources Reuses the shared Vulkan partitioner (`_clone_dim_order`/`expand_copy`/`full`* already registered); WebGPU kernels only. Co-authored-with: Claude Code. ghstack-source-id: 405026073 @exported-using-ghexport Differential Revision: [D111755131](https://our.internmc.facebook.com/intern/diff/D111755131/)
…) to the WebGPU backend Pull Request resolved: #20929 Add `_clone_dim_order` (DMA copy), `aten.expand_copy` (broadcast), and `fill` (`aten.full`/`full_like`/`scalar_tensor`) handlers for the training tail. Key changes: - `runtime/ops/{dim_order,expand_copy,fill}/` — WGSL kernels + handlers (dim_order is a native buffer DMA copy) - `CMakeLists.txt` `WEBGPU_SRCS` — wire the sources Reuses the shared Vulkan partitioner (`_clone_dim_order`/`expand_copy`/`full`* already registered); WebGPU kernels only. Co-authored-with: Claude Code. ghstack-source-id: 405026073 @exported-using-ghexport Differential Revision: [D111755131](https://our.internmc.facebook.com/intern/diff/D111755131/)
…) to the WebGPU backend Pull Request resolved: #20929 Add `_clone_dim_order` (DMA copy), `aten.expand_copy` (broadcast), and `fill` (`aten.full`/`full_like`/`scalar_tensor`) handlers for the training tail. Key changes: - `runtime/ops/{dim_order,expand_copy,fill}/` — WGSL kernels + handlers (dim_order is a native buffer DMA copy) - `CMakeLists.txt` `WEBGPU_SRCS` — wire the sources Reuses the shared Vulkan partitioner (`_clone_dim_order`/`expand_copy`/`full`* already registered); WebGPU kernels only. Co-authored-with: Claude Code. ghstack-source-id: 405026073 @exported-using-ghexport Differential Revision: [D111755131](https://our.internmc.facebook.com/intern/diff/D111755131/)
…) to the WebGPU backend Pull Request resolved: #20929 Add `_clone_dim_order` (DMA copy), `aten.expand_copy` (broadcast), and `fill` (`aten.full`/`full_like`/`scalar_tensor`) handlers for the training tail. Key changes: - `runtime/ops/{dim_order,expand_copy,fill}/` — WGSL kernels + handlers (dim_order is a native buffer DMA copy) - `CMakeLists.txt` `WEBGPU_SRCS` — wire the sources Reuses the shared Vulkan partitioner (`_clone_dim_order`/`expand_copy`/`full`* already registered); WebGPU kernels only. Co-authored-with: Claude Code. ghstack-source-id: 405026073 @exported-using-ghexport Differential Revision: [D111755131](https://our.internmc.facebook.com/intern/diff/D111755131/)
…) to the WebGPU backend Pull Request resolved: #20929 Add `_clone_dim_order` (DMA copy), `aten.expand_copy` (broadcast), and `fill` (`aten.full`/`full_like`/`scalar_tensor`) handlers for the training tail. Key changes: - `runtime/ops/{dim_order,expand_copy,fill}/` — WGSL kernels + handlers (dim_order is a native buffer DMA copy) - `CMakeLists.txt` `WEBGPU_SRCS` — wire the sources Reuses the shared Vulkan partitioner (`_clone_dim_order`/`expand_copy`/`full`* already registered); WebGPU kernels only. Co-authored-with: Claude Code. ghstack-source-id: 405026073 @exported-using-ghexport Differential Revision: [D111755131](https://our.internmc.facebook.com/intern/diff/D111755131/)
…) to the WebGPU backend Pull Request resolved: #20929 Add `_clone_dim_order` (DMA copy), `aten.expand_copy` (broadcast), and `fill` (`aten.full`/`full_like`/`scalar_tensor`) handlers for the training tail. Key changes: - `runtime/ops/{dim_order,expand_copy,fill}/` — WGSL kernels + handlers (dim_order is a native buffer DMA copy) - `CMakeLists.txt` `WEBGPU_SRCS` — wire the sources Reuses the shared Vulkan partitioner (`_clone_dim_order`/`expand_copy`/`full`* already registered); WebGPU kernels only. Co-authored-with: Claude Code. ghstack-source-id: 405026073 @exported-using-ghexport Differential Revision: [D111755131](https://our.internmc.facebook.com/intern/diff/D111755131/)
…) to the WebGPU backend Pull Request resolved: #20929 Add `_clone_dim_order` (DMA copy), `aten.expand_copy` (broadcast), and `fill` (`aten.full`/`full_like`/`scalar_tensor`) handlers for the training tail. Key changes: - `runtime/ops/{dim_order,expand_copy,fill}/` — WGSL kernels + handlers (dim_order is a native buffer DMA copy) - `CMakeLists.txt` `WEBGPU_SRCS` — wire the sources Reuses the shared Vulkan partitioner (`_clone_dim_order`/`expand_copy`/`full`* already registered); WebGPU kernels only. Co-authored-with: Claude Code. ghstack-source-id: 405026073 @exported-using-ghexport Differential Revision: [D111755131](https://our.internmc.facebook.com/intern/diff/D111755131/)
…) to the WebGPU backend Pull Request resolved: #20929 Add `_clone_dim_order` (DMA copy), `aten.expand_copy` (broadcast), and `fill` (`aten.full`/`full_like`/`scalar_tensor`) handlers for the training tail. Key changes: - `runtime/ops/{dim_order,expand_copy,fill}/` — WGSL kernels + handlers (dim_order is a native buffer DMA copy) - `CMakeLists.txt` `WEBGPU_SRCS` — wire the sources Reuses the shared Vulkan partitioner (`_clone_dim_order`/`expand_copy`/`full`* already registered); WebGPU kernels only. Co-authored-with: Claude Code. ghstack-source-id: 405026073 @exported-using-ghexport Differential Revision: [D111755131](https://our.internmc.facebook.com/intern/diff/D111755131/)
Stack from ghstack (oldest at bottom):
Add
_clone_dim_order(DMA copy),aten.expand_copy(broadcast), andfill(aten.full/full_like/scalar_tensor) handlers for the training tail.Key changes:
runtime/ops/{dim_order,expand_copy,fill}/— WGSL kernels + handlers (dim_order is a native buffer DMA copy)CMakeLists.txtWEBGPU_SRCS— wire the sourcesReuses the shared Vulkan partitioner (
_clone_dim_order/expand_copy/full* already registered); WebGPU kernels only.Co-authored-with: Claude Code.
@exported-using-ghexport
Differential Revision: D111755131
Differential Revision: D111755131